Journal of Proteome Research
● American Chemical Society (ACS)
Preprints posted in the last 30 days, ranked by how well they match Journal of Proteome Research's content profile, based on 234 papers previously published here. The average preprint has a 0.16% match score for this journal, so anything above that is already an above-average fit.
Kotli, P. C.
Show abstract
Ancient DNA (aDNA) has transformed the study of hominin relationships, but its preservation in ancient fossils is often limited. Enamel palaeoproteomics offers an alternative molecular approach for taxonomic analysis. In this study, we re-analyse published DDA mass spectrometry data1 from the Denisovan-attributed Penghu 1 mandible (PXD054412)2 and a Neandertal enamel specimen from Gruta de Oliveira, Portugal (PXD038154) 3. Five AMBN peptides carrying the Valine-273 substitution (V273) were validated in the Penghu 1 enamel, three of which were independently detected in both DDA acquisitions. No V273-containing peptide signal was detected in the Neandertal dataset. Conversely, the ancestral Methionine-273 peptide REDPM[+16]AYG was detected exclusively in the Neandertal specimen. Extracted ion chromatograms, isotopic envelope confirmation, and MS2 fragmentation spectra, all support the reported peptide assignments. AMELY-specific peptides additionally support male sex assignment for both ancient individuals. Together, these results confirm the taxon-specific mutual exclusivity of AMBN V273 and M273 variants across Denisovan and Neandertal lineages, establishing AMBN M273V as a molecularly validated diagnostic marker for Denisovan identification from dental enamel. More broadly, targeted MS1 reanalysis of public proteomics datasets provides a scalable complement to ancient genomics for resolving hominin lineage identity and sex determination when DNA is not preserved
Fields, L.; Hubecky, E. M.; Selby, K. G.; Li, L.
Show abstract
Data-independent acquisition (DIA) mass spectrometry has emerged as a powerful tool for neuropeptidomics, but its success relies heavily on the quality of spectral libraries used for peptide identification. There are inherent challenges to mass spectrometry analysis of crustacean neuropeptides, including the endogenous nature in which they are analyzed, extensive post-translational modification (PTM), and atypical fragmentation patterns. Thus, general-purpose proteomic spectral prediction tools may not perform optimally in the endogenous peptide domain. In this study, we benchmark four widely used spectral prediction platforms, Prosit, MS2PIP, AlphaPeptDeep, and UniSpec, to evaluate their performance in predicting the fragmentation of neuropeptides. Using an empirically derived spectral library from crustacean tissues as reference, we assess model compatibility, dot-product similarity, Pearson correlation, and DIA-based identifications across brain, sinus gland, and pericardial organ samples. Our results reveal that no single model comprehensively captures neuropeptide fragmentation characteristics. While UniSpec showed unexpected strengths due to its inclusion of neutral loss ions, AlphaPeptDeep demonstrated the highest spectral similarity, and MS2PIP and Prosit outperformed in DIA-NN identifications. We further highlight the critical impact of neutral loss fragments, present in over 50% of empirical spectra, and emphasize the need for hybrid spectral libraries that integrate complementary strengths across models. This work provides a foundational framework for optimizing spectral library selection in neuropeptidomics and underscores the importance of model-specific biases when analyzing structurally diverse endogenous peptides.
Vasylieva, V.; Massignani, E.; Claeys, T.; Bourassa, F.; Leblanc, S.; Arefiev, I.; Martens, L.; Brunet, M. A.
Show abstract
ShortThe SwissProt database contains a stable 20,418 human protein-coding genes and 42,541 human protein sequences. Ribo-Seq suggests about 7,000 additional, non-canonical Open Reading Frames (ORFs) are present in humans, though only a few of them are confirmed by Mass Spectrometry (MS). Detecting these proteins requires extensive database searches, increasing computational load and inflating False Discovery Rates (FDR). Using the ionbot search engine with the OpenProt database allows for reliable detection of non-canonical proteins while controlling FDR. Ionbot surpasses the Trans-Proteomics Pipeline (TPP) in reproducibility, identifying more peptides and proteins supported by multiple spectra. In addition, open modification searches yield better PSMs compared to closed searches. This work highlights the importance of employing cutting-edge search engines in non-canonical protein research, as well as the value of open modification search in correcting errors in non-canonical protein detection. LongO_ST_ABSBackgroundC_ST_ABSThe SwissProt database reports a quite stable 20,418 human protein-coding genes and 42,541 human protein sequences, figures that have remained stable. New techniques like Ribo-Seq indicate that approximately 7,000 additional, non-canonical Open Reading Frames (ORFs) are translated in humans, few of which have been confirmed by Mass Spectrometry (MS). Detecting these non-canonical proteins requires comprehensive database searches, which increase computational load and False Discovery Rate (FDR). Here, we use the open search engine ionbot in combination with the OpenProt proteogenomics database to reproducibly detect non-canonical proteins while maintaining a well-controlled FDR. ResultsCompared to the current gold standard, the Trans-Proteomics Pipeline (TPP), ionbot shows higher reproducibility, with a higher number of peptides and proteins supported by multiple spectra, and across multiple samples. We observe that PSMs from the open modification search against OpenProt have higher fragment ion intensity correlation compared to PSMs obtained from the closed search, or by only searching canonical proteins. ConclusionsIn this work, we show the potential for open modification searching to correct potential mistakes in non-canonical proteins detection by preventing modified canonical peptides or variants from being incorrectly identified as non-canonical peptides. We also highlight the importance of assessing the FDR of non-canonical identifications separately from canonical ones, as global FDR calculations are biased by the scarcity of non-canonical identifications in each dataset.
de Almeida, R. F.; Fernandes, M.; de Godoy, L. M. F.
Show abstract
The processes such as DNA replication, transcription, and repair are often modulated by specific nuclear proteins, protein-protein interactions (PPIs), and post-translational modifications (PTMs). In Trypanosoma cruzi, however, the nuclear proteome and interactome have not been systematically mapped, limiting the interpretation of nuclear regulatory processes. Here, we report a nuclear proteome resource generated from intact nuclei isolated from T. cruzi and analyzed by high-resolution Orbitrap LC-MS/MS, integrating proteome profiling, computational interaction network inference, and exploratory crosslinking mass spectrometry (XL-MS). Proteome profiling identified 1,734 proteins in the nuclear fraction, including 316 proteins identified with PTM-containing peptides. Subcellular localization prediction and Gene Ontology analysis support nuclear enrichment and highlight functions related to transcription, RNA metabolism, and genome maintenance. The in silico interaction network derived from STRINGDB organizes the proteins into functional clusters, including a histone-associated interaction neighborhood. In parallel, XL-MS identified 26 residue-resolved interprotein crosslinks involving 36 proteins and detected PTMs at or near linked residues. Together, these data support reuse for comparative nuclear proteomics, multi-omics integration, and prioritization of candidates for future functional studies.
Haueis, J. R. S.; Lazar, I. M.
Show abstract
Mass spectrometry (MS) is the leading technology for identifying proteins in complex biological samples. It relies on the use of tandem MS alongside a reference database of canonical protein sequences to computationally identify peptides and their parent proteins. The canonical sequences represent the most widely expressed and functionally validated forms of proteins. Consequently, disease-induced or disease-supportive variants, such as those associated with cancer, will evade detection if they are absent from the database. To address this challenge, this study introduces a revised release of the Unkown Mutation Analysis (XMAn) database by incorporating coding missense and nonsense mutations from the latest versions (v103) of the COSMIC Genome Screen Mutants (GSM) and Cancer Gene Census (CGC) datasets in two distinct FASTA-formatted peptide databases comprising 3,848,499 and 312,658 variants, respectively. The mutated peptides were matched to reviewed, non-redundant UniProt Homo sapiens protein entries (18,362 and 746), and characterized in terms of nucleotide- and amino acid mutation frequencies, peptide length distributions, and associations between specific single-nucleotide (SNV) and single amino acid (SAAVs) variants. Applied to the analysis of MDA-MB-231 breast cancer cell-membrane protein fractions, the database enabled the identification of 300+ high-quality variant peptides - several localized to functional protein-binding and catalytic domains - and 23 aberrant protein products mapped to the CGC dataset. The database is hosted and available for download on Zenodo (XMAn/gsm doi: 10.5281/zenodo.21781023; XMAn/cgc doi: 10.5281/zenodo.21781514) or can be accessed through https://sites.google.com/vt.edu/xman-db/home.
Sato, H.; Akioka, S.; Konno, R.; Okuda, Y.; Ohara, O.; Kawashima, Y.
Show abstract
Serum proteomics is increasingly used for minimally invasive biomarker discovery and disease phenotyping, and the choice of serum preprocessing workflow can shape proteome depth, quantitative characteristics, and downstream biological readouts. However, disease-oriented comparisons within a single cohort remain limited. Here, we compared four serum preprocessing workflows--Top14 depletion (TOP14D), tomato lectin affinity purification (TomAP), and two nanoparticle-based enrichment workflows (NPA and NPB)--using serum from six patients with systemic juvenile idiopathic arthritis (sJIA) and six age- and sex-matched healthy controls, and analyzed them using unified data-independent acquisition mass spectrometry (DIA-MS) and a statistical pipeline. We evaluated proteome depth, missingness, quantitative characteristics, group separation, differential abundance signatures, pathway enrichment, curated sJIA-related gene set coverage, pre-ranked gene set enrichment analysis (GSEA) results, and detection of inflammasome/interferon-related proteins. TomAP yielded the greatest proteome depth (7612 proteins), followed by NPB (6735 proteins) and NPA (6602 proteins), whereas TOP14D yielded the smallest protein set (3303 proteins). Principal component analysis (PCA) showed a separation between the sJIA and control groups for all workflows. Differentially expressed proteins (DEPs) showed limited overlap, with only 75 DEPs common to all four workflows. Functional enrichment patterns were workflow-dependent; TOP14D and TomAP mainly captured neutrophil/myeloid and inflammatory processes, whereas NPA and NPB captured RNA processing- and translation-related signals. TomAP showed relatively broad coverage and positive enrichment of curated sJIA-related gene sets associated with inflammation, innate immunity, and macrophage activation syndrome (MAS). Inflammasome/interferon-related proteins, including NLRC4, PYCARD, GSDMD, MEFV, IL-18, OAS3, and MYD88, showed workflow-dependent detectability and differential abundance. These findings support a disease-oriented benchmark for fit-for-purpose workflow selection according to the disease axis and analytical objective rather than proteome depth alone.
Zang, L.; Grandke, J.; Richter, J.; Kielkowski, P.
Show abstract
Mass spectrometry-based chemical proteomics is a powerful method to analyze proteins labelled by small molecules to identify protein targets of active compounds and to profile protein post-translational modifications. The throughput and high protein input for chemical proteomics workflows has been often a limiting factor for application of the technology for specialized and difficult to culture cell lines. The high protein input was necessary to gain significant difference of noise to signal ratio in proteomics readout. Here, we describe a general chemical proteomics workflow, which is performed in 96-well plate and necessitate only 25 g of protein input to profile post-translationally modified proteins including abundant O-GlcNAcylated proteins as well as low abundant AMPylated proteins. The workflow integrates advances in Cu(I)-catalyzed azide-alkyne cycloaddition to minimize chemical side-reactivity of the click reaction and data-independent acquisition mode during LC-MS/MS measurement. An iterative optimization of protein clean-up on carboxylate-coated paramagnetic beads led to significant saving of the beads usage and lowers the unspecific protein background that resulted in sensitivity gain.
Arauz-Garofalo, G.; Ciordia, S.; Gonzalez de Peredo, A.; Chaoui, K.; Rijal, J. B.; Gaxotte, V.; Folch-i-Casanovas, I.; Azkargorta, M.; Almey, R.; Aloria, K.; Kirim, B. A.; Barderas, R.; Braga-Lagache, S.; Calvo, E.; Chicano-Galvez, E.; Clemente, F.; Chiritoiu, G.; Chiva, C.; Decourcelle, M.; Dhaenens, M.; Diaz, R.; Douche, T.; Duran-Cortines, A.; Duran-Ruiz, M. C.; El Koulali, K.; Escobar-Nino, A.; Fernandez Acero, F. J.; Fernandez-Irigoyen, J.; Garcia-Garcia, C.; Gil, C.; Goetze, S.; Gonzalez Vidal, E.; Gutierrez, M.; Hernaez, M. L.; Lopez, C. M.; Marin-Vicente, C.; Mateos-Martin, M. L.; Mato
Show abstract
Multicenter studies are essential for benchmarking analytical workflows, yet their interpretation is often confounded by the combined effects of experimental protocols and instrumentation. To address this challenge, we introduce a simple normalization-based analytical framework, the recovery metric ({rho}), designed to decouple protocol driven effects from instrument dependent variability. We applied this framework to the 13th Proteomics Multicentric Experiment (PME13), a large multicentric proteomics dataset generated across 27 laboratories using high sensitivity workflows and varying sample preparation protocols. By leveraging a common digested reference sample, {rho} enables direct cross-comparison of all datasets on a unified scale, effectively minimizing instrument-related biases. Using this approach, we demonstrate that apparent instrument dependent trends are largely removed when evaluated through {rho}, revealing consistent protocol driven effects across laboratories. Statistical modeling identified key variables influencing {rho}, including sample input amount, reduction and alkylation, and the use of n-dodecyl-{beta}-D-maltoside (DDM). While DDM was associated with improved {rho}, reduction and alkylation and additional handling steps led to reduced performance, particularly at low input levels. We further highlight practical considerations for the application of ratio based normalization, including the occurrence of values exceeding theoretical bounds, which reflect deviations from underlying assumptions and require appropriate filtering. Overall, this work establishes a generalizable analytical strategy for disentangling confounding factors in multicentric datasets and provides practical guidelines for optimizing high sensitivity proteomics (HSP) workflows. The proposed framework is broadly applicable to other analytical fields where cross laboratory comparability is required.
Antony, F.; Bhattacharya, A.; Aoki, H.; Babu, M.; Duong van Hoa, F.
Show abstract
Quantitative membrane proteomics remains fundamentally limited by sample preparation because detergent extraction can perturb membrane protein interactions, ligand-responsive conformations, and higher-order assemblies before mass spectrometric analysis. Here, we demonstrate that peptide-based surfactants (Peptergents) enable a complete detergent-free workflow for native membrane proteomics. Membrane proteins are extracted directly from biological membranes while preserving their structural and functional integrity and remaining fully compatible with downstream LC-MS/MS workflows. Functional preservation is evidenced by maintenance of ligand-responsive conformations in the ABC transporter MsbA and the endogenous GPCR P2RY12, together with stabilization of the detergent-sensitive nine-subunit holo-translocon HTL, indicating that fragile membrane protein assemblies remain intact. At the proteome level, despite recovering fewer membrane proteins than conventional detergent extraction, Peptergent consistently generates higher peptide signal intensities, retains tissue-specific membrane proteome signatures, and preferentially enriches endoplasmic reticulum-associated metabolic networks, including cytochrome P450 enzymes and their interaction network. Together, these findings establish Peptergents as a broadly applicable membrane extraction technology for LC-MS/MS-based membrane proteomics, preserving native membrane organization and expanding the proteomics toolbox for biochemical, structural, and systems-level analyses of membrane proteins. In Brief StatementThis study establishes Peptergents as a detergent-free membrane extraction technology for LC-MS/MS-based membrane proteomics. Peptergent extraction preserves ligand-responsive membrane proteins, fragile membrane protein assemblies, and tissue-specific membrane proteome signatures while remaining fully compatible with quantitative proteomic workflows. These findings provide a broadly applicable strategy for preserving native membrane organization for biochemical, structural, and systems-level analyses of membrane proteins. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=199 SRC="FIGDIR/small/744532v1_ufig1.gif" ALT="Figure 1"> View larger version (56K): org.highwire.dtl.DTLVardef@1fe34b0org.highwire.dtl.DTLVardef@35400corg.highwire.dtl.DTLVardef@1ffe97aorg.highwire.dtl.DTLVardef@394fc4_HPS_FORMAT_FIGEXP M_FIG C_FIG HighlightsO_LIPeptergents preserve ligand-responsive membrane proteins. C_LIO_LISupport chemoproteomics in thermal proteome profiling assays. C_LIO_LISimplify membrane proteomics workflow. C_LIO_LIMaintain native tissue-specific membrane biology. C_LIO_LIPreserve fragile membrane protein assemblies. C_LI
McDonnell, K.; Geiszler, D. J.; Wamsley, N.; Derks, J.; Sipe, S.; Cohen, Z. A.; Warinner, L. K.; Yeh, M.; Koo, E.; Leduc, A.; Zwang, T. J.; Specht, H.; Slavov, N.
Show abstract
Parallelization of data acquisition substantially increases the throughput of mass spectrometry-based proteomics. However, parallelization also increases the density of mass spectra and consequently the overlap between ions, frustrating their analysis. To improve sequence identification and quantification from such spectra, we developed an open-source software for Joint Modeling of mass spectra (JMod). JMod models overlapping peaks as linear superpositions of their components in both MS1 and MS2 space, which permits multiplexed DIA with smaller mass offsets to increase the multiplexing capacity and thus proteomics throughput for a given plexDIA tag. This enables 9-plexDIA using 2 Da offset PSMtags, increasing throughput 9-fold while preserving quantitative accuracy and coverage depth. Furthermore, we use JMod to deconvolve simultaneous labeling by mass tags and heavy amino acids, thus increasing the throughput of metabolic pulse experiments measuring protein synthesis and degradation rates in single cells from mouse liver. By supporting enhanced decoding of highly multiplexed DIA spectra, JMod provides an open and flexible software that increases the throughput of sensitive proteomics.
Kore, H.; Gleeson, J.; Yin Wan, C.; De Paoli-Iseppi, R.; Dutt, M.; Prawer, Y. D. J.; Alkaraki, A.; Lonsdale, A.; Wells, C.; Smith, L.; Clark, M.; Parker, B.
Show abstract
Quantifying the diversity of RNAs and proteins produced by cells is fundamental to the biological and clinical sciences. However, many RNAs and proteins remain uncharacterised, especially proteins translated from alternate RNA isoforms; untranslated regions of mRNAs and non-coding RNAs, as well as the effects of DNA variation on protein sequences. Proteogenomics aims to characterise the complete proteome by integrating genomics and/or transcriptomics with proteomics, but current tools have limitations in useability, analysis features and visualisation of resulting data. To address these gaps, we developed GenomeProt, a user-friendly GUI-based tool for integrative proteogenomic analysis. We demonstrate its utility by integrating long-read RNA sequencing with mass-spectrometry-based proteomics to pinpoint proteoform expression generated by alternative splicing; discover novel, unannotated proteins in human brain samples; and quantify variant-containing peptides associated with treatment resistance in a melanoma xenograft model. GenomeProt brings the discovery power of proteogenomics to biologists, illuminating the hidden proteome.
Criscuolo, L.; Elmkvist, S. B.; Nawrocki, A.; Jakobsen, L. A.; Jensen, P.; Jensen, P. T.; Huang, H.; Havelund, J. F.; Faergeman, N. J.; Palmisano, G.; Bogetofte, H.; Larsen, M. R.
Show abstract
Comprehensive characterization of protein abundance and multiple post-translational modifications (PTMs) from the same biological samples is essential for understanding cellular regulation and PTM crosstalk but remains analytically challenging. Here, we present STEP-PTM (Sequential Tag-based Enrichment of Post-Translational Modifications), a modular TMT-multiplexed workflow that enables integrated quantitative analysis of the proteome, metabolome and multiple PTM classes from a single peptide preparation. Proteins are digested, isobarically labeled using tandem mass tags (TMT), and combined into a single multiplexed peptide pool prior to sequential PTM enrichment, thereby minimizing technical variability, reducing sample requirements and facilitating direct quantitative integration across datasets. STEP-PTM supports flexible sequential enrichment of phosphopeptides, peptides containing free and reversibly modified cysteines, sialylated N-linked glycopeptides, lysine-acetylated peptides and S-palmitoylated peptides, while preserving non-modified peptides for global proteome analysis. PTM-specific database searches further improve identification confidence and quantitative accuracy, and the modular workflow can readily be adapted by incorporating or omitting enrichment modules according to the biological question. Application of STEP-PTM to TMT16-plex cerebral brain organoids enabled the quantification of 10,413 proteins, 2,969 metabolites, 19,655 phosphopeptides, 28,876 peptides containing reversibly modified cysteines, 9,723 peptides containing free cysteines, 1,716 intact sialylated N-linked glycopeptides and 771 lysine-acetylated peptides from the same biological samples. We further demonstrate the applicability of the workflow to multiple mouse tissues, highlighting its broad utility for integrated systems-level characterization of protein expression and PTM regulation across diverse biological models.
Gerber, Z.; Simard, S.; Kolipaka, H.; Drouin, Z.; Sevigny, J.; Pourcel, V.; del Carmen Crespo Oliva, C.; Tate, B.; Mouzakitis, K.; Placet, M.; Jean, D.; Deuel, K.; Pavlatos, E.; Sturgill, E.; Pucilowska, J.; Mills, G. B.; Labrie, M.
Show abstract
Spatially resolved single-cell proteomic imaging technologies, including cyclic immunofluorescence (CycIF), generate high-dimensional data, critical for tissue-scale biological analysis. However, single-cell analysis remains computationally demanding, lacks standardization across platforms and is often inaccessible to experimental biologists without programming expertise. Here we present SCORPy (Single-Cell proteOmics Research Platform), a standalone, cross-platform desktop application that provides an end-to-end, code-free workflow for the analysis of single-cell proteomic data extracted from imaging experiments. SCORPy introduces methodological advances for preprocessing multiplexed imaging data: an exposure-aware, cycle-matched background correction strategy, and a normalization framework that harmonizes signal distributions across markers while enabling batch correction across experiments. These approaches are integrated with quality control, interactive thresholding and cell phenotyping using a hierarchical cell reference library, and downstream compositional and spatial analyses within a unified interface. Sample-level metadata can be incorporated throughout the workflow to support integrative analyses and facilitate generation of publication-ready visualizations. By combining robust preprocessing methods with an accessible implementation, SCORPy reduces computational barriers and promotes broader adoption of spatial single-cell proteomics analysis.
Meijer, M.; Hong, J.; Pohl, T.; Koudelka, T.; Bassot, C.; Hoernberg, H.; Lee, S.; Rho, H. S.; Lee, A. C.; Pelechano, V.; Piazza, I.
Show abstract
Spatial proteomics aims to resolve protein composition within intact tissues, yet extraction-based liquid chromatography-mass spectrometry (LC-MS) workflows face an inherent trade-off: smaller sampling units increase spatial specificity, whereas larger sampling units provide greater proteome depth and robustness. As analytical sensitivity improves, sampling-unit size therefore becomes a key experimental design parameter. Current extraction-based LC-MS workflows typically rely on laser capture microdissection (LCM), where sample recovery and scalability can become limiting at low input. Spatially resolved laser-activated cell sorting (SLACS) offers an alternative tissue-isolation strategy based on single-pulse near-infrared laser activation. Here, we use SLACS to systematically examine the resolution-sensitivity trade-off across sampling units ranging from single-cell-equivalent to larger low-input tissue regions. Few-cell sampling retained substantial proteomic information relative to larger regions while increasing spatial specificity. Applied to the mouse somatosensory cortex, SLACS generated deep, layer-resolved proteomic profiles from regions corresponding to approximately 60 cells and preserved major layer-specific molecular patterns at inputs as low as approximately 6 cells. These results highlight sampling-unit size as an important experimental design parameter in extraction-based spatial proteomics and support few-cell sampling as a practical compromise between spatial specificity, proteome depth and robustness.
Zakar-Polyak, E.; Kerepesi, C.
Show abstract
Contextualized protein-protein interaction networks provide crucial insight into diseases and other biological processes, but for a profound understanding of such processes and their distinct effects on individuals, the protein-protein interactions within individual samples must be investigated. A straightforward approach to estimate the PPI network of a sample is to restrict a general network of known PPIs to the proteins that are found in the sample. Although proteomics methods are becoming more accessible and precise, large-scale and single-cell studies still mainly target characterizing the transcriptomics profile of the samples, which is then often used as an approximation of the protein activities. The correlation of gene expression and protein abundance has been addressed in the past, but information about the deviations of the different omics-based estimates of the PPI networks is still lacking. In this study, we performed a comparative analysis of transcriptomic-based and proteomic-based sample-specific PPI network estimates to fill this gap. We created a framework for a comprehensive and transparent comparison of the two omics levels in two independent datasets, with a special focus on time-related network dynamics. We found that the size-adjusted characteristics of the different omics-based networks are very similar; the overall trend of how they change with time is also often the same, but the rate of the changes typically differs. The characteristics of the nodes present in both types of networks also show high similarity and often different time-related rates of change, but this varies among metrics. These results shed light on the properties of PPI network estimations and advise caution in interpreting them appropriately.
Ni, J.; Tracey, H.; Hao, L.
Show abstract
Stem cells secrete diverse extracellular proteins that regulate pluripotency, differentiation, and cell-cell communication, making them powerful model systems for studying development, disease mechanisms, and regenerative medicine. However, robust stem cell secretome analysis remains technically challenging. Unlike many other cell types, stem cells cannot tolerate serum starvation or growth factor deprivation, while low-abundance secreted proteins are often masked by media-derived proteins and intracellular contamination. Here, we systematically optimized the secretome proteomics workflow in iPSCs, by evaluating culture medium composition, conditioned-media collection time, cell plating density, media harvest and preparation methods, LC-MS acquisition methods, and data analysis strategies. Full-strength Essential 8 medium, 48 h media collection, 80% cell confluency, two-step centrifugation, and data-independent acquisition (DIA)-LC-MS/MS provided the optimal secretome proteomics data quality. We then applied the optimized platform to an isogenic iPSC disease model to investigate how progranulin deficiency reshapes the extracellular and intracellular proteomes. Progranulin-deficient iPSCs showed a coordinated reduction of extracellular lysosomal hydrolases despite relatively modest intracellular proteome changes, suggesting altered lysosome trafficking and possible impairment of lysosomal exocytosis. Together, this work establishes a robust and standardized workflow for stem cell secretome proteomics and demonstrates its utility for investigating extracellular proteome remodeling in human disease models.
Handelmann, C.; Miles, A. K.; Ye, Y.; Freire, M.; Dewhirst, F. E.; Chen, T.; Mark Welch, J.; Kauffman, K. M.
Show abstract
Metaproteomics aims to capture a taxonomically comprehensive snapshot of proteins in a sample. Design of reference databases is a key aspect of metaproteomic workflows, as these define what is ultimately seen. Databases tailored to focal biomes offer optimal performance, yet their construction often requires drawing on heterogeneous data sources, posing a challenge to reproducibility and documentation. Here we present maniFasta, a tool enabling users to generate standardized, reproducible, and robustly documented protein reference sets from diverse input sources and datatypes. Users provide information on their desired input types and sources, and the output is an integrated database comprising a protein sequence file (FASTA), with harmonized identifiers and standardized headers, and an associated provenance metadata table (manifest). We highlight the value of maniFasta in the context of salivary metaproteomics, addressing the need for a taxonomically comprehensive reference database. The AllOralsDB resource includes human proteins, as well as proteins from bacteria and archaea, fungi and other microeukaryotes, viruses and viroid-like elements, dietary sources, and common contaminants. Together, this work provides a community resource for oral and salivary metaproteomics (https://www.homd.org/ftp/AllOralsDB/), and a versatile and accessible tool for constructing protein databases for metaproteomics generally (https://github.com/KauffmanLab/maniFasta).
Veth, T. S.; Sutherland, E.; Hinkle, J. D.; Bergen, D.; Melani, R. D.; McAlister, G. C.; Mullen, C.; Riley, N. M.
Show abstract
Glycan heterogeneity is a fundamental property of glycoproteins. A holistic understanding of glycan modification states is critical to translating glycoproteome regulation to biological function, but the high degree of glycosite-level heterogeneity leads to technical challenges in measuring glycoproteoforms. Common bottom-up glycoproteomics provide some insights but cannot recapitulate the full ensemble of glycoproteoforms from glycopeptide measurements alone. Promising efforts to profile masses of intact glycoproteins have recently explored data-independent acquisition (DIA) coupled with proton-transfer charge reduction (PTCR) or electron-capture-induced charge reduction mass spectrometry (MS). While valuable for generating broad glycoproteoform mass distributions, these approaches have remained limited in their ability to generate discrete glycoproteoform mass measurements, largely because they rely on low-resolving power measurements and deconvolution that does not account for isotopic information. Here, we develop a DIA-PTCR workflow that couples high-resolving power (Rp ~240,000 at m/z 200) tandem mass spectra with an open-source processing suite to define glycoproteoform populations within 20 ppm mass accuracy thresholds. We demonstrate the glycoproteoform characterization capabilities of this platform using a collection of glycoproteins with well-described translational interests (EpCAM, TIGIT, CD40, PDL1, and CD24). With a focus on EpCAM, we showcase how intact glycoproteoform masses acquired using our high-resolving power DIA-PTCR (hRp-DIA-PTCR) approach can be integrated with bottom-up intact glycoproteomics and Direct-Mass Technology (i.e., Orbitrap-based charge-detection MS) acquisitions to inform structural and biological insights. Altogether, our hRp-DIA-PTCR method extends the current capabilities of intact glycoprotein analyses by enabling robust characterization of isotopically resolved proteoforms and facilitating deep biological interpretation of glycosylation heterogeneity. Our open-source informatics platform includes a GUI-based tool called PTsliCR to clean PTCR spectra directly from DIA-PTCR raw files and a deconvolution R package called IsoTrac, both of which are freely available on GitHub at https://github.com/riley-research.
Diaz, A.; Tichshenko, N.; Depoortere, B. G. J.; Andrade Buono, R.; De Geest, P.; Vranken, W. F.; Martens, L.; Ramasamy, P.
Show abstract
Post-translational modifications (PTMs) and genetic variants regulate protein function, signalling, and disease, but their interpretation requires integration of sequence annotations with structural, interaction, and biophysical context. Although resources such as Scop3P, UniProt, the Protein Data Bank, and AlphaFold provide extensive annotations and structural information, integrating these data into reproducible structure-aware analyses still requires custom scripting and manual coordination between multiple independent tools. To address this challenge, we developed Scop3P-Toolkit, an open-source executable analytical environment for interactive analysis of PTMs, mutations, and proteomics-derived peptides in their structural context. The toolkit integrates protein annotation retrieval with structural mapping, residue interaction network analysis, comparative structural analysis, and residue-level biophysical profiling within a unified framework. Experimentally supported phosphosites, phosphopeptides, and phosphoproteomics evidence are provided for human proteins through Scop3P, with optional integration of curated UniProt PTM annotations. UniProt-derived PTMs, sequence features, and genetic variants are available for proteins from any species, extending the framework beyond the human phosphoproteome. Scop3P-Toolkit supports structure-centric analyses including interpretation of PTMs and disease-associated variants, analysis of residue interaction networks and their rewiring across alternative conformations, structural localisation of peptides, and exploration of protein-protein, protein-ligand, and host-pathogen interfaces. Interactive visualisation links sequence annotations, three-dimensional structures, residue interaction networks, and biophysical profiles, enabling coordinated exploration across multiple molecular representations. The toolkit is distributed as Jupyter notebooks, browser-based Voila applications, and a Galaxy interactive tool, providing transparent, accessible, and reproducible workflows for both computational and experimental researchers. By integrating biological annotation resources into executable, structure-aware workflows, Scop3P-Toolkit enables reproducible interpretation of PTMs, mutations, and proteomics data.
Pusparum, M.; Thas, O.; Ertaylan, G.
Show abstract
Conventional univariate reference intervals (UniRIs) are widely used to identify abnormal biomarker values, but they evaluate each biomarker independently and do not account for coordinated deviations between biomarkers. We developed and evaluated a joint reference region (JRR) framework for plasma proteomics data using the Olink proteomics dataset generated by the UK Biobank Pharma Proteomics Project, covering approximately 3,000 plasma proteins. JRRs were estimated for selected protein pairs in a healthy reference subset, while UniRIs were estimated separately for individual proteins using the nonparametric method. Both approaches were then evaluated in ICD-defined disease subsets. Biomarker discovery revealed sparse and heterogeneous disease--protein associations, with some proteins recurring across multiple phenotypes and others showing more disease-specific patterns. The added value of JRRs varied across diseases and protein pairs. Across evaluated protein pairs, 56.5\% showed higher sensitivity under the JRR framework than the UniRI of the first protein, and 47.3\% showed higher sensitivity than the UniRI of the second protein. At the disease level, the median proportion of protein pairs with improved JRR sensitivity was 0.57. JRRs were most informative when univariate detection was limited but a subset of diseased observations was flagged only by the joint region. These findings suggest that JRRs provide a complementary approach to UniRIs by capturing abnormal joint biomarker configurations in high-dimensional proteomics data.